Papers with domain expert mixture
DEMix Layers: Disentangling Domains for Modular Language Modeling (2022.naacl-main)
Copied to clipboard
| Challenge: | Extensive experiments with autoregressive transformer LMs show that DEMix layers reduce test-time perplexity and increase training efficiency. |
| Approach: | They introduce a new domain expert mixture layer that enables conditioning a language model on the domain of the input text. |
| Outcome: | Experiments with 1.3B LMs show that DEMix layers reduce test-time perplexity, increase training efficiency, and enable rapid adaptation. |